Journal: Scientific Reports
Article Title: A multimodal comparison of latent denoising diffusion probabilistic models and generative adversarial networks for medical image synthesis
doi: 10.1038/s41598-023-39278-0
Figure Lengend Snippet: Illustration of the Medfusion model. ( A ) General overview of the architecture. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$x$$\end{document} x and \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\widetilde{x}$$\end{document} x ~ are the input and output images. ( B ) Details of the autoencoder with a sampling of the latent space via the reparameterization trick at the end of the encoder and a direct connection (dashed lines) into the decoder (only active for training the autoencoder). ( C ) Detailed view of the denoising UNet with a linear layer for time and label embedding. ( D ) Detailed view of the submodules inside the autoencoder and UNet. If not specified otherwise, a convolution kernel size of 3 × 3, GroupNorm with 8 groups, and Swish activation was used.
Article Snippet: Therefore, we performed an additional experiment: Figure 1 Reconstruction quality of Medfusion Variational Autoencoder (VAE).
Techniques: Sampling, Activation Assay